This is the benchmark dataset 1 described in the manuscript.

Step1_input-pacbio-ccs-read.fasta: Raw PacBio CCS reads of 10 known ~1,5 kb MSG plamid mixture.
Step1_output_Step2_input-uclust-cluster.uc: Uclust clustering results. Read names are assigned to each cluster group. The actual read sequences of each group can be retrieved from above file for setp 3.
Step2_output_Step3_input-consense-sequence.fasta: The consensus sequenced derived from step2 and it is also the input file of step 3.
Step3_output-final-contig.fasta: The final contig file derived froms step3 - Sequencher cleanup step.


Expected_10_MSG_control_seq.fasta: The expected answer for the benchmark dataset. To evaluate the effectivness of the described analytical pipeline, you can compare the sequences between this file and Step3_output-final-contig.fasta